NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

LoCoRe: Image Re-Ranking with Long-Context Sequence Modeling

https://doi.org/10.1109/CVPR52734.2025.00895

Xiao, Zilin; Suma, Pavel; Sachdeva, Ayush; Wang, Hao-Jen; Kordopatis-Zilos, Giorgos; Tolias, Giorgos; Ordonez, Vicente (June 2025, IEEE Computer Vision and Pattern Recognition (CVPR))

Free, publicly-accessible full text available June 10, 2026
ViC-MAE: Self-supervised Representation Learning from Images and Video with Contrastive Masked Autoencoders

Hernandez, Jefferson; Villegas, Ruben; Ordonez, Vicente (September 2024, European Conference on Computer Vision (ECCV), Springer, Cham)

We propose ViC-MAE, a model that combines both Masked AutoEncoders (MAE) and contrastive learning. ViC-MAE is trained using a global representation obtained by pooling the local features learned under an MAE reconstruction loss and using this representation under a contrastive objective across images and video frames. We show that visual representations learned under ViC-MAE generalize well to video and image classification tasks. Particularly, ViC-MAE obtains state-of-the-art transfer learning performance from video to images on Imagenet-1k compared to the recently proposed OmniMAE by achieving a top-1 accuracy of 86% (+1.3% absolute improvement) when trained on the same data and 87.1% (+2.4% absolute improvement) when training on extra data. At the same time, ViC-MAE outperforms most other methods on video benchmarks by obtaining 75.9% top-1 accuracy on the challenging Something something-v2 video benchmark. When training on videos and images from diverse datasets, our method maintains a balanced transfer-learning performance between video and image classification benchmarks, coming only as a close second to the best-supervised method.
more » « less
Full Text Available
PropTest: Automatic Property Testing for Improved Visual Programming

https://doi.org/10.18653/v1/2024.findings-emnlp.483

Koo, Jaywon; Yang, Ziyan; Cascante-Bonilla, Paola; Ray, Baishakhi; Ordonez, Vicente (November 2024, Findings of the Association for Computational Linguistics)

Full Text Available
ViC-MAE: Self-supervised Representation Learning from Images and Video with Contrastive Masked Autoencoders

Hernandez, Jefferson; Villegas, Ruben; Ordonez, Vicente (September 2024, European Conference on Computer Vision (ECCV))

Full Text Available
Grounding Language Models for Visual Entity Recognition

Xiao, Zilin; Gong, Ming; Cascante-Bonilla, Paola; Zhang, Xingyao; Wu, Jie; Ordonez, Vicente (September 2024, European Conference on Computer Vision (ECCV))

Full Text Available
ElasticDiffusion: Training-Free Arbitrary Size Image Generation Through Global-Local Content Separation

https://doi.org/10.1109/CVPR52733.2024.00631

Haji-Ali, Moayed; Balakrishnan, Guha; Ordonez, Vicente (June 2024, IEEE Conference on Computer Vision and Pattern Recognition (CVPR))

Full Text Available
Improved Visual Grounding through Self-Consistent Explanations

https://doi.org/10.1109/CVPR52733.2024.01244

He, Ruozhen; Cascante-Bonilla, Paola; Yang, Ziyan; Berg, Alexander C; Ordonez, Vicente (June 2024, IEEE Conference on Computer Vision and Pattern Recognition (CVPR))

Full Text Available
SCoRD: Subject-Conditional Relation Detection with Text-Augmented Data

https://doi.org/10.1109/WACV57701.2024.00563

Yang, Ziyan; Kafle, Kushal; Lin, Zhe; Cohen, Scott; Ding, Zhihong; Ordonez, Vicente (January 2024, IEEE)

Full Text Available
Improving Visual Grounding by Encouraging Consistent Gradient-Based Explanations

https://doi.org/10.1109/CVPR52729.2023.01837

Yang, Ziyan; Kafle, Kushal; Dernoncourt, Franck; Ordonez, Vicente (June 2023, IEEE Conference on Computer Vision and Pattern Recognition)

Full Text Available
Estimating and Maximizing Mutual Information for Knowledge Distillation

Shrivastava, Aman; Qi, Yanjun; Ordonez, Vicente (January 2023, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops)

In this work, we propose Mutual Information Maximization Knowledge Distillation (MIMKD). Our method uses a contrastive objective to simultaneously estimate and maximize a lower bound on the mutual information of local and global feature representations between a teacher and a student network. We demonstrate through extensive experiments that this can be used to improve the performance of low capacity models by transferring knowledge from more performant but computationally expensive models. This can be used to produce better models that can be run on devices with low computational resources. Our method is flexible, we can distill knowledge from teachers with arbitrary network architectures to arbitrary student networks. Our empirical results show that MIMKD outperforms competing approaches across a wide range of student-teacher pairs with different capacities, with different architectures, and when student networks are with extremely low capacity. We are able to obtain 74.55% accuracy on CIFAR100 with a ShufflenetV2 from a baseline accuracy of 69.8% by distilling knowledge from ResNet-50. On Imagenet we improve a ResNet-18 network from 68.88% to 70.32% accuracy (1.44%+) using a ResNet-34 teacher network.
more » « less
Full Text Available

« Prev Next »

Search for: All records